Back

Journal of Medical Imaging

SPIE-Intl Soc Optical Eng

Preprints posted in the last 90 days, ranked by how well they match Journal of Medical Imaging's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.

2026-06-29 radiology and imaging 10.64898/2026.06.26.26356692 medRxiv
Top 0.1%
6.7%
Show abstract

Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.

2
Standardised evaluation and monitoring of site-specific AI performance with physical CT phantoms

Genske, U.; Laudani, A.; Yan, L.; Peng, Y.; Boening, G.; Ulas, S. T.; Wagner, M. P.; Diekhoff, T.; Hamm, B.; Jahnke, P.

2026-07-02 radiology and imaging 10.64898/2026.07.01.26357033 medRxiv
Top 0.1%
4.9%
Show abstract

Artificial intelligence (AI) applications in computed tomography (CT) imaging require objective and continuous testing, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised testing and monitoring of AI, demonstrated in liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that anatomically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.

3
Deep Learning for Automated Meningioma Segmentation: Toward Clinical Integration and Workflow Efficiency

Fenney, E.; Muralidharan, L.; Ruffle, J. K.; Pandit, A.; Millip, M.; Hammam, A.; Brookes, T.; Jabeen, F.; Colman, J.; Sarwani, O.; Alattar, K.; Efthymiou, E.; Kallam, N.; Siddiqui, J.; Marcus, H. J.; Nachev, P.; Hyare, H.

2026-05-15 neurology 10.64898/2026.05.12.26352585 medRxiv
Top 0.1%
4.0%
Show abstract

Background: Meningiomas are the most common primary intracranial tumors in adults, and volumetric assessment increasingly guides surveillance and treatment decisions. Automated segmentation could enable standardized volumetry but requires robust validation. Purpose: To develop a fully automated three-dimensional deep learning model for meningioma segmentation on multiparametric MRI, and to evaluate segmentation accuracy, external generalizability, failure modes, radiologist-rated clinical plausibility, and workflow feasibility. Methods: From 2024 to 2026, this retrospective study trained a custom 3D nnU-Net residual encoder model. Expert segmentations covered enhancing tumor (ET), tumor core (TC), and whole tumor (WT). Dice similarity coefficient (DSC) was the primary metric. External validation used an independent single-institution dataset (n = 310 intracranial cases) with incomplete MRI protocols. Failure modes, model equity, and inference time were assessed. A blinded multi-rater study (10 radiologists; 510 cases) rated TC segmentations using a 0-10 Likert scale, analyzed with linear mixed-effects models. Results: Model training used the BraTS Meningioma 2023 dataset (n = 1000; mean age 60.2 {+/-} 14.5; 705 female). In cross-validation, mean DSC was 0.939 for ET, 0.937 for TC, and 0.921 for WT. In external validation, mean DSC was 0.872 for TC and 0.842 for WT, despite heterogeneous protocols and incomplete sequences. Predicted TC volumes correlated strongly with reference volumes in cross-validation (r = 0.995) and external validation (r = 0.971). Most common failure modes were skull base and intraosseous tumors with performance equitable across demographic subgroups. Mean inference time was 1.2 seconds. In blinded evaluation (1120 ratings), model segmentations received higher scores than reference annotations (+0.32 BraTS; +1.38 external validation). Conclusion: A fully automated deep-learning model achieved high meningioma segmentation accuracy across multi-institutional training data and external clinical imaging. In a blinded study, model segmentation quality exceeded reference annotations, and 1.2-second inference supported workflow integration. Prospective evaluation is warranted before routine deployment.

4
Retrieval-Augmented Claude Opus 4.7 and GPT-5.5 Surpass Human Performance on the Nuclear Cardiology Board Preparation Exam (and Claude Drafts a Paper About it)

Killekar, A.; Shanbhag, A.; Miller, R. J.; Dey, D.; Bourque, J.; Phillips, L.; Chareonthaitawee, P.; Slomka, P.

2026-05-13 radiology and imaging 10.64898/2026.05.08.26352768 medRxiv
Top 0.1%
3.6%
Show abstract

BackgroundPrevious studies evaluated large language model (LLM) performance on the American Society of Nuclear Cardiology (ASNC) Board Preparation Exam. Without domain-specific context, the best model (GPT-4o) achieved 63.1%, below the estimated 65% passing threshold and the 78% mean score of human fellows-in-training (FITs). Providing textbook context improved GPT-4o to 73.8% on text-only questions, but still fell short of human trainees. Whether next-generation LLMs with retrieval-augmented generation (RAG) can exceed this gap is unknown. MethodsClaude Opus 4.7 and GPT-5.5 were administered all 168 questions (141 text-only, 27 image-based) from the 2023 ASNC Board Preparation Exam across 5 iterations each, using RAG with a nuclear cardiology textbook, companion atlas, and ASNC clinical guidelines. Claude used local FAISS-based semantic retrieval; GPT-5.5 used Azures cloud-hosted vector store. Performance was compared to prior LLM results and 13 human FITs. ResultsAcross 5 iterations, Claude Opus 4.7 achieved a mean accuracy of 86.3% {+/-} 1.4% (text 88.8%, image 73.3%). GPT-5.5 achieved 86.7% {+/-} 2.2% (text 88.5%, image 77.0%) but refused a mean of 12.2 questions (7.3%) per iteration due to safety filters. Both models surpassed the human FIT mean (78.0%) and the estimated passing threshold. Compared to GPT-4o without context (63.1%), this represents a 23-percentage-point improvement in 18 months. ConclusionNext-generation LLMs with RAG now surpass average human trainee performance on nuclear cardiology board preparation questions, suggesting significant potential as educational tools and knowledge-reference aids in cardiovascular imaging. Condensed AbstractAcross 5 iterations each, Claude Opus 4.7 and GPT-5.5 with retrieval-augmented generation achieved mean accuracies of 86.3% and 86.7% on the 2023 ASNC Board Preparation Exam (168 questions), both surpassing the mean human fellow-in-training score of 78%. GPT-5.5 refused a mean of 12.2 questions (7.3%) per iteration due to safety filters. These results represent a 23-percentage-point improvement over the best prior LLM without context (63.1%), demonstrating that RAG-enhanced LLMs have reached human-level proficiency in nuclear cardiology knowledge. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/26352768v2_ufig1.gif" ALT="Figure 1"> View larger version (49K): org.highwire.dtl.DTLVardef@5f2465org.highwire.dtl.DTLVardef@4e80d3org.highwire.dtl.DTLVardef@1ebbb93org.highwire.dtl.DTLVardef@167d3c1_HPS_FORMAT_FIGEXP M_FIG C_FIG Overview of the three-study research arc evaluating LLM performance on the 2023 ASNC Board Preparation Exam. Study 1 (2024) tested four LLMs without context (best: GPT-4o, 63.1%). Study 2 (2025) added textbook context to GPT-4o (73.8%). Study 3 (2026, current) evaluated Claude Opus 4.7 and GPT-5.5 with retrieval-augmented generation across 5 iterations each (mean 86.3% and 86.7%, respectively), both surpassing the human fellow-in-training mean of 78%. Right panel shows the performance scale with key thresholds.

5
An Automated CT-derived Marker of Renal Tumor Complexity: The CLARITY Score

Jonnalagadda, R.; Patel, S. H.; Abusafieh, H. T.; Seshadri, R.; Jevnikar, D.; Younis, S.; Al-Bayati, A.; Saputro, N.; Knorr, J.; Wang, B.; Ozery-Flato, M.; Rosen-Zvi, M.; Abouassaly, R.; Remer, E.; Heller, N.; Weight, C.

2026-05-12 urology 10.64898/2026.05.08.26352647 medRxiv
Top 0.1%
3.4%
Show abstract

Background and ObjectiveSurgical complexity for renal tumors has traditionally been assessed using manual nephrometry scores, which require unreimbursed physician effort and are subject to interobserver variability. This study introduces an objective, fully automated alternative derived from decades of experience at a large academic center. MethodsWe trained a CT classification model to predict whether a patient would ultimately undergo Partial or Radical Nephrectomy (PN or RN). We hypothesized that the models confidence in RN (termed the CLARITY score) would serve as a surrogate for the difficulty of nephron-sparing approaches and thus for tumor complexity. This hypothesis was tested using multivariate logistic regression for failure to achieve trifecta, estimated blood loss (EBL) [&ge;] 500 mL, and length of stay [&ge;] 3 d. CLARITY was compared with tumor size and R.E.N.A.L. score. External validation in a geographically distinct cohort was performed. Key Findings and LimitationsFor predicting RN, CLARITY achieved an AUROC of 0.899 internally and 0.898 externally. In the external PN subgroup, it outperformed tumor size and R.E.N.A.L. score in predicting failure to achieve trifecta (AUROC 0.613), EBL [&ge;] 500 mL (0.727), and length of stay [&ge;] 3 d (0.673). In multivariable analysis, CLARITY remained associated with each outcome, whereas R.E.N.A.L. and size were not. This study is limited by its retrospective design. Conclusions and Clinical ImplicationsCLARITY is an automated CT-derived marker that quantifies renal tumor complexity more effectively than tumor size and R.E.N.A.L. score and may support scalable, objective preoperative complexity assessment. To support reproducibility and external validation, we have released a public inference pipeline and web-based DICOM upload portal for research use.

6
An Open, Reproducible Gamma-Variate Pipeline for CT-Perfusion Time-Attenuation Curve Analysis, with Standardized (ASIST-Japan) Map Visualization

Yamamoto, S.

2026-06-29 radiology and imaging 10.64898/2026.06.26.26356666 medRxiv
Top 0.1%
3.2%
Show abstract

CT perfusion (CTP) is central to acute-stroke and oncologic imaging, yet quantitative outputs vary substantially across vendor software, undermining reproducibility. We present an open, transparent core (ctp-core) that fits first-pass time-attenuation curves with a gamma-variate model, derives perfusion indices (peak enhancement, time-to-peak, bolus-arrival time, and area under the curve) analytically from the fitted parameters, and renders parametric maps with the ASIST-Japan standardized lookup table (a-LUT) so that visualization is comparable across sites. Every parameter, bound, and processing step is exposed. The method is validated on Monte-Carlo synthetic curves with known ground truth; no confidential or patient data are used. Across signal-to-noise ratio (SNR) levels 5 to 100 (200 independent runs per level) the pipeline recovers peak time to within 0.03-0.52 s and peak amplitude to within 0.4-8.1% (mean absolute error), degrading monotonically with noise; at a representative SNR of 20 it recovers peak time within 0.13 s, peak amplitude within 2.0%, and bolus-arrival time within 0.51 s, with fit quality R-squared = 0.98. The reproducibility demonstration is deterministic (fixed seed) and re-runs to bit-stable metrics. All code, the synthetic-data generator, the standardized-visualization module, evaluation scripts, and a 34-test suite are released openly for independent verification. The contribution is a fully open, parameter-transparent gamma-variate plus standardized-visualization pipeline with reproducible synthetic benchmarks: a reference others can audit, reuse, and build on.

7
TCIA Radiology Image Processing for AI and Radiomics

Rich, J. M.; Kang, R.; Jin, D.; Subramanian, S.; Duddalwar, V.; Pachter, L.

2026-06-24 radiology and imaging 10.64898/2026.06.15.26354651 medRxiv
Top 0.1%
3.2%
Show abstract

We developed a standardized, reproducible preprocessing framework for computed tomography (CT) imaging data from multi-institutional repositories such The Cancer Imaging Archive (TCIA), enabling consistent radiomics and artificial intelligence (AI) analyses. Imaging data from TCGA-KIRC patients available on TCIA were used as a representative heterogeneous dataset characterized by variation in acquisition protocols, inconsistent metadata, and differing image quality. The proposed modular pipeline includes series filtering, DICOM-to-NIfTI conversion, orientation harmonization to a canonical coordinate system, voxel spacing normalization, intensity clipping and normalization, segmentation integration, and metadata validation, and is implemented in a reproducible, notebook-based framework compatible with common radiomics and deep learning workflows. This pipeline standardizes imaging data into analysis-ready volumes with consistent geometry, intensity distributions, and spatial alignment, reducing non-biological variability that can adversely affect radiomic feature stability and model performance. The modular design enables task-specific adaptation of individual preprocessing steps while maintaining overall consistency. Although demonstrated on TCIA, this framework is generalizable to other heterogeneous imaging datasets and provides a foundation for robust, large-scale computational imaging studies.

8
Fine-Tuning SAM2 for Coronary Artery Segmentation in X-Ray Fluoroscopy

Sivakumar, E.

2026-06-19 radiology and imaging 10.64898/2026.06.16.26355803 medRxiv
Top 0.1%
3.2%
Show abstract

SAM2 (Meta, 2024) provides a strong starting point for segmentation, but given the unique challenges in medical imaging (noise from patient movement, the projection-based nature of X-ray fluoroscopy, and low contrast between vessels and background), direct application is difficult. We fine-tune MedSAM2 on annotated coronary angiograms and apply it to video data for point-of-care use. On the ARCADE validation set (200 images), the fine-tuned model achieves Dice 0.767 compared to 0.033 zero-shot. On 10 fluoroscopic video studies from CoronaryDominance, it tracks vessels coherently and avoids falsely segmenting ribs, stents, and bypass grafts in 9 of 10 studies. Code is available at https://github.com/elakiyasivakumar/SAM2-Coronary-Angiography-VA and the fine-tuned checkpoint at https://huggingface.co/Elakiya17/CA-SAM2.

9
Feature-Based Parametric Response Mapping on Thoracic Computed Tomography for Robust Disease Classification in COPD

Namvar, A.; Shan, B.; Hoff, B.; Labaki, W. W.; Murray, S.; Bell, A. J.; Galban, S.; Kazerooni, E. A.; Martinez, F. J.; Hatt, C. R.; Han, M. K.; Galban, C. J.; Ram, S.

2026-04-27 radiology and imaging 10.64898/2026.04.24.26351675 medRxiv
Top 0.1%
2.8%
Show abstract

PurposeTo develop an interpretable feature-based Deep Parametric Response Mapping (PRMD) method that combines wavelet scattering convolution networks and machine learning to spatially detect and quantify functional small airways disease (fSAD) and emphysema on paired inspiratory-expiratory CT scans, with enhanced noise robustness. Materials and MethodsIn this retrospective analysis of prospectively acquired data (2007-2017), we developed and validated a deep learning-based PRM approach using paired CT scans from 8,972 tobacco-exposed COPDGene participants ([&ge;]10 pack-years; mean age 60.1 {+/-} 8.8 years; 46.5% women), including controls with normal spirometry (n = 3,872; controls), PRISm (n = 1,089), GOLD 1-4 COPD (n = 4,011). Data were stratified into training, validation, and testing sets (24:6:70). PRMD extracts translation-invariant image features using a wavelet scattering network and applies a subspace learning classifier to classify voxels as emphysema or non-emphysematous air trapping (fSAD). PRMD was compared with conventional density-based PRM for voxel-wise agreement, correlation with pulmonary function, robustness to noise, and sensitivity to misregistration using Pearson correlation, Bland-Altman analysis, and paired t tests. ResultsPRMD achieved 95% voxel-wise agreement with standard PRM (r = 0.98) while demonstrating significantly greater robustness under noise. PRMD showed stronger correlations with FEV (emphysema: r = -0.54; fSAD: r = -0.51; P < 0.0001) than standard PRM (r = -0.42 for both; P < 0.0001). Under simulated high-noise conditions, standard PRM overestimated disease by [~]15%, whereas PRMD limited error to < 5% (P < 0.001). ConclusionPRMD provides an interpretable, feature-driven and noise-resilient alternative to traditional PRM for emphysema and fSAD classification, enhancing the reliability of CT-based COPD phenotyping for multi-center studies and low-dose imaging applications. Key PointsO_LIThis study introduces combined wavelet scattering and subspace learning for medical image segmentation, enabling accurate, interpretable voxel-level classification of emphysema and functional small airways disease on paired CT scans. C_LIO_LIThe proposed Deep Parametric Response Mapping method demonstrated 95% voxel-wise agreement with standard Parametric Response Mapping and stronger correlations with spirometric measures, enhancing the clinical relevance of CT-based phenotyping for Chronic Obstructive Pulmonary Disease. C_LIO_LIDeep Parametric Response Mapping significantly improved robustness to image noise--reducing overestimation of emphysema and functional small airways disease from [~]15% to <5% (P < 0.001)--and benefits from reduced data requirements due to the fixed, mathematically defined filters used in wavelet scattering. C_LI Summary StatementDeep Parametric Response Mapping improves the accuracy and noise robustness of CT-based classification of emphysema and functional small airways disease using feature-based representations, enhancing the reliability of COPD phenotyping.

10
Failure detection in medical image classification under realistic distribution shifts: A large-scale benchmark

Steinmetz, P.; Frouin, F.; Morard, V.; Buvat, I.

2026-05-05 radiology and imaging 10.64898/2026.05.04.26350496 medRxiv
Top 0.1%
2.4%
Show abstract

Medical images (MI) exhibit variability due to different acquisition protocols, devices, and patient populations, making failure detection at inference time essential for reliable deployment of clinical classifiers. As existing evaluations of failure detection methods use different settings, it is difficult to compare results and identify the best strategy, if any. We present a comprehensive benchmark of eight confidence scoring functions and two score-aggregation strategies across eight MI tasks spanning diverse modalities, backbone architectures, training setups, and failure sources. The confidence ranking ability and classification error mitigation are jointly evaluated. While no single method systematically dominated across settings, aggregation of confidence scores consistently matched or approached the best individual method and substantially reduced silent failure rate. The failure detection performance was strongly correlated with classifier accuracy for all tested settings. These findings provide large-scale evidence regarding the strengths and limitations of confidence scoring strategies and offer actionable guidance for mitigating silent failures under realistic distribution shifts in MI.

11
Scan length as a major driver of CT radiation dose: a diagnostic reference level audit from Kosovo

Rudi, G.; Vula, F.; Bicaku, A.; Dedushi, K.; Ahmetgjekaj, I.

2026-05-17 radiology and imaging 10.64898/2026.05.12.26353024 medRxiv
Top 0.1%
2.2%
Show abstract

Computed tomography is the largest contributor to population radiation dose from medical imaging, yet no diagnostic reference levels (DRLs) have been published from Kosovo or the Western Balkans. This retrospective audit analyzed all CT examinations performed on a 128- slice scanner at the University Clinical Centre of Kosovo between January and March 2026. After exclusions, 1,535 acquisitions from 1,092 patients across nine examination categories were analyzed. Local DRLs were defined as the 75th percentile and compared against German (BfS 2022) and Turkish (Kahraman et al., 2024) reference values. Head CT (n = 590) demonstrated CTDIvol 4.7% below the BfS DRL yet scan length 98.5% above the orientation value (median 25.8 vs 13 cm). Abdomen-pelvis CTDIvol matched the BfS reference while scan length exceeded it by 28%. Coronary CTA showed CTDIvol +377%, consistent with retrospective ECG gating. Excess scan length, not CTDIvol, is the major driver of elevated dose at this institution. The identified excesses are correctable through technologist landmarking training, protocol review, and enabling iterative reconstruction.

12
Dual-Filament 3D Printing of Patient-Specific CT Phantoms with Embedded Implants and Tunable Metal-Artifact Intensity

Pasyar, P.; Mei, K.; Im, J. Y.; Roshkovan, L.; Geagan, M.; Noël, P. B.

2026-07-20 radiology and imaging 10.64898/2026.07.17.26358319 medRxiv
Top 0.1%
2.2%
Show abstract

ABSTRACT Background: Metallic implants such as orthopedic screws, prostheses, and dental hardware produce beam-hardening, photon-starvation, and streak artifacts that degrade computed tomography (CT) image quality, and the metal artifact reduction (MAR) methods developed to mitigate them require objective, reproducible benchmarking. Purpose: Objective evaluation of MAR algorithms in CT is hindered by the absence of phantoms that simultaneously provide anatomically realistic backgrounds, embedded implants of known geometry, and controllable, ground-truth--referenced artifact intensity. We present a dual-filament, voxel-level three-dimensional (3D) printing method that fulfills these requirements and demonstrate its capabilities on a clinically representative cervical spine case with embedded orthopedic spinal screws. Methods: The proposed method extends the PixelPrint framework, a fused-deposition-modeling (FDM) workflow that converts clinical Digital Imaging and Communications in Medicine (DICOM) data directly into 3D-printer Geometric code (G-code) without intermediate segmentation or surface meshing, to interleaved, voxel-level deposition of two filaments: a calcium-doped polylactic acid (PLA) for soft tissue and bone, and a higher-attenuation metal-doped PLA for metallic implants. For demonstration, anonymized DICOM data of a healthy cervical spine were used to design and fabricate three matched phantoms, each with six embedded spinal screws at C4--C6: a 0% metal-infill ground-truth phantom, a 50% medium-metal-infill phantom, and an 85% high-metal-infill phantom. All phantoms were scanned on a clinical spectral CT system at 120 kVp and 1000 mAs, reconstructed at 0.67 mm slice thickness with virtual monoenergetic imaging (VMI) across 50--190 keV. Method performance was characterized by region of interest (ROI)-based Hounsfield Unit (HU) agreement with the source patient data and by the noise-independent Gumbel-distribution p-index metric. Results: The dual-filament method reproduced patient anatomy, soft-tissue contrast, and screw geometry with high fidelity. ROI HU values agreed with patient data within {+/-}25 HU for soft tissue and trabecular bone; cortical regions were underestimated owing to the current ceiling of the calcium-doped PLA used in this study. The tunable-artifact behavior was quantified as follows: the Gumbel location parameter scaled monotonically from 46.7 HU (no-metal background) to 57.1 HU (50% infill) to 90.5 HU (85% infill) for the VMI 70 keV with standard filter. High-keV VMI reconstructions substantially reduced streak and beam-hardening artifacts while preserving anatomic detail. Conclusions: The proposed dual-filament, voxel-level PixelPrint method enables the fabrication of patient-specific, multi-material CT phantoms with embedded metallic implants and controllable, ground-truth--referenced artifact intensity. Although demonstrated here in a single cervical-spine case, the workflow is anatomy- and implant-agnostic by construction and could in principle be adapted to other musculoskeletal sites (e.g., knee, hip, dental) and implant materials, providing a reproducible methodological foundation for benchmarking MAR algorithms, characterizing spectral CT performance, and validating emerging photon-counting detector systems. Keywords: 3D printing methodology; fused deposition modeling; voxel-level multi-material printing; spectral computed tomography; metal artifact reduction; phantom design; orthopedic implants; dual filament; PixelPrint.

13
Analysis and Mitigation of Equipment-induced Shortcuts in AI Models for Laparoscopic Cholecystectomy

Protserov, S.; Repalo, A.; Mashouri, P.; Hunter, J.; Masino, C.; Madani, A.; Brudno, M.

2026-04-24 surgery 10.64898/2026.04.22.26351545 medRxiv
Top 0.1%
2.0%
Show abstract

Machine learning models have seen a lot of success in medical image segmentation domain. However, one of the challenges that they face are confounders or shortcuts: spurious correlations or biases in the training data that affect the resulting models. One example of such confounders for surgical machine learning is the setup of surgical equipment, including tools and lighting. Using the task of identification of safe and dangerous zones of dissection in laparoscopic cholecystectomy images and videos as a use-case, we inspect two equipment-induced biases: the presence of surgical tools in the field of view and the position of lighting. We propose methods for evaluating the severity of these biases and augmentation-based methods for mitigating them. We show that our tool bias mitigations improve the models consistency under tool movements by 9 percentage points in the most inconsistent cases, and by 4 percentage points on average. Our lighting bias mitigations help reduce fraction of true dangerous zone pixels that may be predicted as safe under light changes from 5% to 1.5%, without compromising segmentation quality.

14
TopBrain Segmentation Challenge for Whole Brain Vessel Anatomy

Yang, K.; Shi, P.; Huang, H.; Musio, F.; Baazaoui, H.; Aydin, O. U.; Hilbert, A.; Hamadache, R. E.; Yalcin, C.; Zhang, M.; Falcetta, D.; de la Rosa, E.; Shit, S.; Prabhakar, C.; Wittmann, B.; Rokuss, M. R.; Kirchhoff, Y.; Al-Maskari, R.; Hoeher, L.; Juchler, N.; Casamitjana, A.; Cleary, J.; Schmick, A.; Baumgartner, P.; Deseoe, J.; Vandans, O.; Lee, D.; Oh, K.; LaBella, D.; Mazher, M.; Niederer, S. A.; Qayyum, A.; Liu, Y.; Chen, J.; Kim, W.; Asawalertsak, N.; Kim, M.; Shin, D.; Park, S.-H.; Kikuchi, S.; Zhang, Y.; Liu, J.; Cui, Y.; Qiu, Y.; Verschuur, A.; Zhang, J.; van der Schaaf, I.; Su, R.;

2026-05-30 radiology and imaging 10.64898/2026.05.28.26354312 medRxiv
Top 0.1%
1.9%
Show abstract

We present the TopBrain 2025 Challenge, the first benchmark for fine-grained multiclass segmentation of the whole brain vasculature in both computed tomography angiography (CTA) and magnetic resonance angiography (MRA). Building on the TopCoW challenge, TopBrain scales vessel annotation from the Circle of Willis to the entire brain, introducing a dataset of 90 annotated volumes across 48 landmark vessel classes spanning arterial and venous systems, of which 50 training volumes are publicly released. Vessel definitions were consolidated from established neuroanatomical references into a unified annotation scheme, and vessel caliber measurements along the centerline are reported for the first time across the whole brain vascular anatomy. To address the unique challenges of multiclass brain vessel segmentation, we propose an evaluation framework that accounts for detection in segmentation performance, assesses anatomical plausibility, and introduces novel contamination metrics that characterize inter-class prediction errors. Fifteen teams from over 220 registered participants submitted algorithms to the benchmark. The top-performing teams built on nnUNet with principled system design choices, achieving around 80% Dice scores, near-zero invalid neighbor counts, over 60% F1 scores for side-road vessels, and below 18% foreground contamination ratio. Larger vessels are easier to segment, while smaller and more complex vessels remain the true bottleneck. The annotated datasets and podium-finish algorithms are made publicly available on Zenodo.

15
SortIT - A Tool For Assessing Observer Variability And Creating Ground Truth Image Classification Datasets

Uegami, W.; Bisson, T.; Okoshi, E. N.; Costa da Silva, F. G.; Jiragawasan, C.; Zerbe, N.; Bychkov, A.; Fukuoka, J.

2026-05-29 pathology 10.64898/2026.05.28.728616 medRxiv
Top 0.2%
1.8%
Show abstract

Interobserver variability in pathological assessments is a well-recognized challenge that impacts diagnostic reliability and disease understanding. This variability exists across many subspecialties due to the subjective nature of evaluations. Artificial intelligence (AI) applied to whole slide images has potential to standardize procedures and reduce variability in pathology, but transitioning to these technologies does not guarantee improvement. Establishing reliable ground truth datasets with consensus annotations is crucial for developing robust AI solutions. We introduce SortIT, an open-source web application that facilitates systematic creation and evaluation of ground truth image tile annotations. SortIT enables multiple annotators to independently label tiles, with flexible user permission controls. Annotated data can be exported for statistical analysis of observer variation and for creating ground truth datasets from consensus tiles. We outline protocols using SortIT for several use cases: (1) mitosis segmentation in tumor regions, (2) evaluating AI solutions for prostate cancer grading by comparing to expert consensus, and (3) granuloma classification by annotating discriminative tile-level features. Key strengths of SortIT lies in its ease of deployment, making it accessible and usable for a wide range of users. Overall, SortIT provides a valuable tool to establish high-quality ground truth datasets and comprehensively assess observer variability. Critical evaluation of ground truth quality using systematic annotation methodologies is crucial for developing accurate and generalizable diagnostic AI tools. Its open-source nature facilitates community adoption and further development.

16
Determinants and propagation of velocity uncertainty in 2D phase-contrast MRI

Rodriguez-Soto, A. E.; Schuchardt, E. L.; Narayan, H. K.; Printz, B. F.; Hegde, S.; Hopkins, S. R.; Contijoch, F.

2026-06-04 radiology and imaging 10.64898/2026.06.01.26353730 medRxiv
Top 0.2%
1.7%
Show abstract

Purpose: To quantify the contributions of signal-to-noise ratio (SNR) and velocity-to-encoding ratio (v/VENC) to velocity uncertainty in phase-contrast (PC) MRI and to develop a framework for in vivo voxel-wise uncertainty estimation. Methods: Through-plane 2D PC-MRI of the ascending aorta was acquired using multiple velocity encodings (150, 200, 300 cm/s) and flip angles (0, 5, 15, 20 degrees) to vary v/VENC and SNR. Voxel-wise SNR and velocity uncertainty maps were generated using empirically calibrated phase-noise modeling. Phase-resolved subject-level analyses were performed to quantify the relative contributions of SNR and |v|/VENC to percent velocity uncertainty (%unc). Uncertainty was propagated to flow, stroke volume (SV), and cardiac output (CO). Results: Velocity uncertainty varied substantially across the cardiac cycle and depended on both SNR and |v|/VENC. Across cardiac phases, |v|/VENC accounted for most explained variance in %unc (partial R2=0.666), while SNR provided a smaller but meaningful contribution (partial R2=0.287; full R2=0.909). Near peak systole, SNR contributed more strongly while overall uncertainty remained low. In contrast, diastolic %unc became unstable as velocity approached zero. These effects were most pronounced at low |v|/VENC, where higher VENC settings increased uncertainty despite similar SNR. SV uncertainty ranged from 0.27% to 1.07% across VENCxFA protocols. Conclusion: Velocity uncertainty in PC-MRI depends on both SNR and VENC adequacy in a physiologically phase-dependent manner. Relative uncertainty may become inadequate for precise quantification in low-flow applications, such as diastolic regurgitant jets, despite adequate SNR. Spatiotemporal uncertainty mapping provides a framework for uncertainty-aware PC-MRI acquisition and interpretation.

17
Quantitative Assessment of Dual and Triple Energy Window Scatter Correction in Myocardial Perfusion SPECT with a 4D Phantom

El Bab, M.; Guvenis, A.

2026-04-25 cardiovascular medicine 10.64898/2026.04.17.26351095 medRxiv
Top 0.2%
1.7%
Show abstract

Conflicting evidence on scatter correction (SC) methods plagues quantitative myocardial perfusion SPECT (MPI), hindering standardized clinical protocols. This simulation study, utilizing the SIMIND Monte Carlo program and a highly realistic 4D XCAT phantom, systematically evaluates Dual Energy Window (DEW, with k=0.5) and Triple Energy Window (TEW) SC techniques. We uniquely investigate their performance across various photopeak window widths (2, 4, and 6 keV) and novel overlapped/non-overlapped configurations specifically for Tc 99m MPI - parameters largely unexplored in realistic cardiac models. Images were reconstructed with OSEM under uncorrected (UC), SC, and combined attenuation and scatter corrected (ACSC) conditions. Quantitative analysis focused on signal-to-noise ratio (SNR), contrast-to-noise ratio (CNR), defect contrast, and relative noise to background (RNB). Our findings consistently show ACSCs superior performance in CNR, SNR, and defect contrast, confirming its critical role. Interestingly, SC alone reduced noise but compromised defect contrast relative to UC, highlighting a potential trade-off without attenuation correction. Crucially, this study reveals minimal influence of photopeak window width and overlap configuration on image quality, and no significant difference between DEW and TEW across most metrics. These results provide essential evidence for optimizing quantitative MPI protocols, suggesting that for Tc-99m, the choice between DEW and TEW, and specific window settings, may be less critical than ensuring robust attenuation correction.

18
FreqFuseNet: Resolving Feature-Scale Mismatch in Dual-Frequency Fusion for Thin-Wall Head-and-Neck OAR Segmentation

Chen, W.-Y.; Wan, S.-Y.; Lin, G.-Y.

2026-07-13 radiology and imaging 10.64898/2026.07.09.26357642 medRxiv
Top 0.2%
1.7%
Show abstract

Accurate segmentation of thin-wall organs-at-risk (OARs)-the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear-is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863-fold activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5 percent residual coefficient to behave as an approximately 43-fold dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5 percent residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 million parameters versus 414.6 million for the full wavelet baseline.

19
RadGuide AI: Development and Technical Evaluation of a General Nuclear Medicine Agent for Traceable Radiopharmaceutical Decision Support

Gu, X.; Zhu, H.; Zhong, F.; Teng, G.-J.

2026-07-10 radiology and imaging 10.64898/2026.07.09.26357614 medRxiv
Top 0.2%
1.5%
Show abstract

Background: Nuclear medicine and radiopharmaceutical development require coordinated radiochemistry, dosimetry, molecular imaging, radiation-safety and clinical decision processes. Current workflows remain fragmented, difficult to audit and poorly standardised for evaluating domain-specific AI support. Methods: We developed RadGuide AI, a nuclear medicine agent built around a traceable data-model-tool loop. Patent, literature and clinical-trial records were converted into 15,596 initial QA items; relevance screening, completeness checks, semantic deduplication and cross-validation retained 5,474 core QA items. MedGemma-27B-Instruct served as the foundation model and was adapted with LoRA. The system incorporated 55 MCP-wrapped tools covering radiopharmaceutical R&D, clinical decision support, imaging analysis and radiation-safety/dosimetry. Evaluation used a locked N=200 benchmark with predefined denominators, leakage control, expert scoring, statistical procedures, factuality audits and tool-execution metrics. Results: RadGuide-LLM achieved 88.5% answer accuracy (177/200; 95% CI, 83.3-92.2%) and a Macro-Average score of 21.5/25 (bootstrap 95% CI, 20.9-22.0), exceeding GPT-4o, DeepSeek-V3.2 and the base MedGemma model in this technical evaluation. Supplementary audits reported guideline compliance, terminology recall, knowledge coverage, tool-routing success and preclinical/phantom dosimetry agreement with explicit denominators and confidence intervals. Interpretation: RadGuide AI converts nuclear medicine queries into auditable retrieval, tool selection, calculation, verification and reporting workflows. The findings support technical feasibility, not definitive patient-level clinical validation; prospective multicentre studies and external benchmark release remain required before clinical deployment.

20
MRI-Based Pressure Gradient Mapping in Patient-Specific Models of Coarctation of the Aorta

Nair, P.; Ferrari, L.; Loecher, M.; McGrath, C. M.; Castillo Passi, C. A.; Marsden, A. L.; Ennis, D. B.

2026-06-03 radiology and imaging 10.64898/2026.05.27.26353898 medRxiv
Top 0.2%
1.5%
Show abstract

Purpose: Accurate assessment of the pressure gradient ({Delta}P) across aortic coarctation (CoA) is critical for determining disease severity and the need for intervention. Current non-invasive methods are unreliable, while invasive catheterization remains the clinical gold standard. This study evaluates a novel MRI acquisition strategy, 4D-FlowP, that simultaneously encodes blood velocity and acceleration to enable reliable non-invasive pressure gradient mapping in CoA. Methods: Patient-specific compliant aortic phantoms were created from clinical MRI data of two patients with CoA. Additional geometries were synthetically generated by increasing stenosis severity. Phantoms were studied in an MRI compatible flow loop under physiologically realistic flow and pressure conditions. Pressure gradients were estimated using conventional 4D-Flow MRI, 4D-FlowP, and fluid-structure interaction (FSI) simulations. Results were compared against ground-truth catheter-based measurements across multiple flow rates and stenosis severities. Results: Conventional 4D-Flow consistently underestimated {Delta}P (slope = 0.63, R2=0.75) relative to catheter measurements. In contrast, 4D-FlowP demonstrated substantially improved agreement (slope = 0.95, R2=0.75). FSI simulations showed the highest overall agreement with catheter-derived {Delta}P (slope = 1.14, R2=0.82). Scan times for 4D-FlowP were comparable to 4D-Flow (26 vs. 24 minutes). Conclusion: 4D-FlowP enables a more accurate MRI-based pressure gradient mapping in CoA than conventional 4D-Flow, when compared to ground truth catheter measurements. These findings support further in vivo evaluation of 4D-FlowP as a non-invasive alternative for functional assessment of CoA severity